Papers with visual question answering
Copied to clipboard
| Challenge: | Neural Module Networks are a class of neural networks that involve human-specified neural modules . current models only learn the parameters of the modules and/or the order of their execution . |
| Approach: | They propose to learn internal structure and sequence without extra supervisory signals . they use dynamically composable modules which are then assembled into a layout . |
| Outcome: | The proposed model performs comparable to models using hand-designed modules. |
Copied to clipboard
| Challenge: | Existing frameworks for Integrative AI lack flexibility and composability to handle multimodal tasks. |
| Approach: | They propose a configurable framework for Integrative AI that orchestrates multiple pre-trained models to conduct complex multimodal tasks. |
| Outcome: | The proposed framework achieves impressive results on zero-shot multimodal tasks . it can communicate and personalize for users, and it can be used in a multimodal agent . |
Copied to clipboard
| Challenge: | a new open-source library for language-vision research and applications is available for free. |
| Approach: | They introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications. |
| Outcome: | The proposed library is open-source and highly extensible and configurable. |
Copied to clipboard
| Challenge: | Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment. |
| Approach: | This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding. |
| Outcome: | This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog. |
Copied to clipboard
| Challenge: | Standardized tests have been proposed as replacements to the Turing test as a driver for progress in AI. |
| Approach: | et al. propose standardized tests as replacements to the Turing test as a driver for progress in AI. |
| Outcome: | a series of standardized tests have been proposed as replacements to the Turing test . the tutorial categorizes open domain and closed domain tests into two categories . open domain tests require the system to have significant domain knowledge and reasoning capabilities. |
Copied to clipboard
| Challenge: | In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Approach: | This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Outcome: | This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
Copied to clipboard
| Challenge: | Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored. |
| Approach: | They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
| Outcome: | The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
Copied to clipboard
| Challenge: | FlagEvalMM is an evaluation framework designed to assess multimodal models . it is designed to be used for vision-language understanding and generation tasks . |
| Approach: | They propose an evaluation framework that decouples model inference from evaluation through an independent evaluation service. |
| Outcome: | The evaluation framework offers accurate and efficient insights into model strengths and limitations. |
Copied to clipboard
| Challenge: | Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets. |
| Approach: | They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining . |
| Outcome: | This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area. |
Copied to clipboard
| Challenge: | Existing systems for visual question answering are overfitted to training data and are sensitive to small perturbations. |
| Approach: | They propose a robustness measure to augment visual question answering datasets to measure generalization capabilities. |
| Outcome: | The proposed model can quantify failure cases which reveal that current systems are still brittle. |
Copied to clipboard
| Challenge: | Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing. |
| Approach: | They propose a retrieval-augmented multi-modal transformer architecture for embedding images and captions in the same space. |
| Outcome: | The proposed approach improves visual question answering over strong baselines and hot-swapping indices. |
Copied to clipboard
| Challenge: | Existing evaluation datasets feature Western-centric images and English text, while their non-English counterparts are often derived from the latter. |
| Approach: | They propose to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. |
| Outcome: | The proposed model underperforms in visual understanding and dialect-specific generation across four Arabic-speaking countries. |
Copied to clipboard
| Challenge: | Lack of perceptual grounding limits vision-language models' ability to interpret visual data . prior work on visualized data understanding focused on adapting VLMs to instruction tuning and chain-of-thought supervision . |
| Approach: | They propose a framework that enhances visual reasoning through human-like interpretation grounding. |
| Outcome: | The proposed framework improves on ChartQA and ChartQAPro benchmarks by +11.2%. |
Copied to clipboard
| Challenge: | Existing approaches to visual question answering (VQA) are not suitable for real-world applications. |
| Approach: | They propose a supervised multi-modal domain adaptation method for visual question answering in images that exploits supervised domain adaptation. |
| Outcome: | The proposed method outperforms state-of-the-art methods on the benchmark VQA 2.0 and VizWiz datasets. |
Copied to clipboard
| Challenge: | Multi-modal Large language models still suffer from model hallucination and lack of specific knowledge when answering challenging questions. |
| Approach: | They propose to use a multi-modal retrieval augmented generation method to integrate knowledge from all modalities into a model to enable alignment between query and knowledge. |
| Outcome: | The proposed method achieves significant performance improvement on the VQA dataset. |
Copied to clipboard
| Challenge: | a new wave of large vision–language models (LVLMs) incorporate images as input in addition to text . a recent study examined potential gender and racial biases in such systems based on the perceived characteristics of the people in the input images. |
| Approach: | They examine potential gender and racial biases in large vision–language models . they query a dataset of AI-generated images of people to see whether they differ . |
| Outcome: | The proposed dataset shows that the images differ in gender and race according to the perceived characteristics of the person depicted. |
Copied to clipboard
| Challenge: | Large multi-modal models (LMMs) are revolutionizing the way machines interact with the world, unlocking new possibilities across multi-dimensional applications. |
| Approach: | They propose a parameter-efficient fine-tuning strategy that combines both . they find that parameter tuning methods distort the feature representation space . |
| Outcome: | The proposed strategy preserves representation space while limiting performance on downstream tasks. |
Copied to clipboard
| Challenge: | Recent approaches for developing vision and language models leverage existing vision and a language expert and try to learn a mapping between them. |
| Approach: | They propose to use a resampler module to create a ‘visual prompt’ which is then fed to the large language models (LLM) using a textual prompt. |
| Outcome: | The proposed method has been shown to be effective across coarse-grained tasks like image captioning and visual question answering, but more fine-grounded tasks that require spatial understanding have not been thoroughly examined. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) tasks are often based on directives, which can cause ambiguities in human utterances. |
| Approach: | They propose a method that clarifies ambiguous questions using gaze information . they propose combining gaze information with gaze information to improve accuracy . |
| Outcome: | The proposed method improves performance in some cases of a GazeVQA system on Gaze. |
Copied to clipboard
| Challenge: | Despite its versatility, CLIP-based applications often suffer from misunderstandings regarding user intent, leading to discrepancies between the required number of objects and the actual outputs. |
| Approach: | They empirically evaluate CLIP’s understanding of quantity from text, image, and cross-modal perspectives by carefully designing different experimental settings and datasets. |
| Outcome: | The proposed model has shown significant success in various downstream tasks, including editing, generation, and quality evaluation. |
Copied to clipboard
| Challenge: | a new tool for evaluating expressive cross-modal interactions is needed . empirical multimodally-additive function projection is a tool for isolating unimodal structure . |
| Approach: | They propose a tool that modifies model predictions so that cross-modal interactions are eliminated . they propose to use EMAP to evaluate models' ability to leverage cross-module interactions . |
| Outcome: | The proposed tool can be used to evaluate models on image+text classification tasks . it finds that removing cross-modal interactions results in little to no performance degradation . |
Copied to clipboard
| Challenge: | a framework for visual question answering is based on modular code generation . the scope of reasoning needed for visual questions is vast, and requires many skills . |
| Approach: | They propose a framework that formulates visual question answering as modular code generation. |
| Outcome: | The proposed framework improves accuracy on COVR and GQA datasets by 3% and 2% compared to the few-shot baseline that does not employ code generation. |
Copied to clipboard
| Challenge: | Existing VQA benchmarks focus on factual correctness but rarely capture what information users actually find useful. |
| Approach: | They propose a framework to quantify how much information an image–question pair provides . they conduct experiments with several state-of-the-art VLMs to determine their reliability . |
| Outcome: | The proposed framework quantifies how much information an image–question pair provides in hospitality contexts. |
Copied to clipboard
| Challenge: | Existing benchmarking datasets for visual question answering focus on machine "understanding" but it remains unclear how progress on those datasets corresponds to improvements in this real-world use case. |
| Approach: | They evaluate the visual question answering task by evaluating a variety of VQA models. |
| Outcome: | The proposed model can achieve high scores on tasks thought to require human-like comprehension, including image tagging and captioning. |
Copied to clipboard
| Challenge: | a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks . |
| Approach: | They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities . |
| Outcome: | The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors. |
Copied to clipboard
| Challenge: | Existing techniques for visual question answering focus on English questions, but many applications require a multilingual module. |
| Approach: | They propose a deep learning framework for multilingual and code- mixed visual question answering . they create Hindi and Code-mixed VQA datasets by exploiting linguistic properties of these languages . |
| Outcome: | The proposed model is capable of predicting answers from the questions in Hindi, English or Code- mixed (Hindi-English) languages. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a task of answering open-ended questions about images. |
| Approach: | They evaluate two vision-and-language (V&L) models under different settings . they find they tend to learn to solve the benchmark rather than the skills required by VQA . |
| Outcome: | The proposed models exhibit poor generalization under out-of-distribution settings. |
Copied to clipboard
| Challenge: | Existing tasks to generate question-answer pairs from visual images are under-explored. |
| Approach: | They propose a task that targets question-answer pair generation from visual images. |
| Outcome: | The proposed model can generate diverse or consistent QAPs on two benchmarks. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) have unlocked many complex use cases that require Multi-Modal (MM) understanding and MM generation. |
| Approach: | They propose a plug-and-play technique that adds relevant retrieved information to prompts as few-shot examples during inference. |
| Outcome: | The proposed method significantly improves the output quality of large vision language models when input prompts are augmented with relevant information retrieved by Vision-Language retrievers like UniRAG. |
Copied to clipboard
| Challenge: | Large visionlanguage models (LVLMs) are a powerful visual-language reasoning tool. |
| Approach: | They propose to integrate attention analysis with LLaVA-CAM to determine interactions between visual representations. |
| Outcome: | The proposed approach can be used to determine interactions between visual representations. |
Copied to clipboard
| Challenge: | Using vision language models, we examine demographic biases in VLMs across gender, race, age, and skin tone. |
| Approach: | They propose a benchmark for uncovering demographic biases in Vision Language Models . they propose 'Gras Bias Score' to quantify bias in VLMs based on gender, race, age and skin tone . |
| Outcome: | The proposed model achieves 98, far from the unbiased ideal of 0. |
Copied to clipboard
| Challenge: | Current explanation generation models are trained to select the best answers from Multiple-Choice questions or to classify single-word answers to a predetermined vocabulary. |
| Approach: | They propose a multitask learning approach towards a Unified Model for Answer and Explanation generation (UMAE) UMAE models surpass the prior state-of-the-art answer accuracy on A-OKVQA by 10 15%, show competitive results on OK-VQA and VCR, and demonstrate promising out-of domain performance on VQA-X. |
| Outcome: | The proposed model outperforms the state-of-the-art model on A-OKVQA and VCR and shows promising out-of domain performance on VQA-X. |
Copied to clipboard
| Challenge: | Existing methods for training language-vision models only consider monolingual learning, especially English. |
| Approach: | They propose to extend an English language-vision model into a multilingual and code-mixed model by using knowledge distillation techniques. |
| Outcome: | The proposed model outperforms existing models on multilingual and code-mixed VQA datasets on eleven languages. |
Copied to clipboard
| Challenge: | Vision-language models struggle on culturally situated inputs, study shows . despite impressive performance, many VLMs struggle on such culturally grounded inputs . |
| Approach: | They propose a new margin-based selector to identify neurons associated with cultural selectivity . they also introduce a model-dependent decoder to identify such neurons . |
| Outcome: | The proposed model outperforms probability- and entropy-based methods in identifying neurons associated with cultural selectivity. |
Copied to clipboard
| Challenge: | Object detection is used in vision and language tasks but is expensive to learn . popular models rely on annotating ground-truths for bounding boxes and semantic labels . empirically, object detection leads to effective transfer learning and improved captioning and visual question answering models . |
| Approach: | They examine the effect of decoupling box proposal and featurization on down-stream tasks . they propose a family of "two-stage" object detectors that propose category-agnostic bounding boxes . |
| Outcome: | The proposed method improves image captioning and visual question answering models by leveraging large amounts of labeled annotations. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a multimodal machine learning problem that challenges a model to answer a question posed about an image. |
| Approach: | They propose a generative model enhanced by multimodal prompt retrieval that integrates retrieved prompts and multimodal features to generate answers in free text. |
| Outcome: | The proposed model outperforms its non-retrieval counterpart by 30% on medical VQA tasks. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) often hallucinate and produce captions that mention concepts that cannot be found in the image. |
| Approach: | They propose to add grounding objectives to captions that explicitly align image regions or objects to text spans to reduce hallucination. |
| Outcome: | The proposed evaluation protocol reduces the amount of hallucination in LVLMs by adding grounding objectives. |
Copied to clipboard
| Challenge: | Adapting visual programming to specialized tasks or domains remains challenging due to high annotation and inference costs. |
| Approach: | They propose a low-cost visual program distillation method that can be used for models with at most 1 billion parameters and requires no human-generated program annotations. |
| Outcome: | The proposed method can generate high-quality visual programs with no human-generated annotations with a relatively small amount of question/answer data. |
Copied to clipboard
| Challenge: | Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts. |
| Approach: | They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset. |
| Outcome: | The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages. |
Copied to clipboard
| Challenge: | Chain-of-thought (CoT) prompting is a prompting strategy that improves reasoning in large language models, but its effectiveness in vision-language models remains limited due to over-reliance on textual cues and memorized knowledge. |
| Approach: | They propose a visual question-answering dataset derived from driving theory exams that incorporates textual explanations with visual tokens extracted from entities relevant to the reasoning process. |
| Outcome: | The proposed approach outperforms chain-of-thought prompting in large language models and vision-language models in real-world scenarios. |
Copied to clipboard
| Challenge: | Existing tasks in Visual Entity Linking (VEL) rely on textual data to complement multi-modal linking or only link objects with general entities. |
| Approach: | They propose a task to link regions of images with corresponding entities in Knowledge Bases . they propose three sub-tasks, based on a human-annotated visual person dataset . |
| Outcome: | The proposed task is based on a human-annotated visual person linking dataset . the proposed sub-tasks are validated on the WIKIPerson dataset based upon the proposed methods . |
Copied to clipboard
| Challenge: | Existing research addresses ambiguous visual questions by rephrasing questions, but it fails to address the inherently interactive nature of user interactions with visual language models (VLMs). Existing studies focus on re-phrase questions, and lack of a benchmark to assess VLMs’ capacity for resolving ambiguities through interaction. |
| Approach: | They propose a visual question answering task that provides a natural language answer to a question based on a given image and an automated pipeline to generate ambiguity-clarification question pairs. |
| Outcome: | The proposed benchmark targets three common categories of ambiguity in visual question answering (VQA) context and encompasses various VQA scenarios. |
Copied to clipboard
| Challenge: | Large pre-trained models have proved to be remarkable zero- and (prompt-based) few-shot learners in unimodal vision and language tasks. |
| Approach: | They propose to use frozen unimodal models to learn a lightweight mapping between the representation spaces of unimod models using aligned image-text data. |
| Outcome: | The proposed method can generalize to unseen VL tasks from a few in-context examples while training orders of magnitude fewer parameters. |
Copied to clipboard
| Challenge: | NegVQA is a visual question answering (VQA) benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions. |
| Approach: | They propose a visual question answering benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions. |
| Outcome: | The proposed model fails to correctly interpret negation, leading to critical errors in interactive AI systems. |
Copied to clipboard
| Challenge: | Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning. |
| Approach: | They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks. |
| Outcome: | The proposed model types do not consistently improve self-rationalization in multimodal tasks. |
Copied to clipboard
| Challenge: | Multimodal large language models combine visual and textual data for tasks like image captioning and visual question answering. |
| Approach: | They propose temperature scaling and iterative prompt optimization to calibrate MLLMs and enhance model reliability. |
| Outcome: | The proposed techniques improve MLLMs and improve model reliability. |
Copied to clipboard
| Challenge: | Visual dialog (VisDial) requires a dialog agent to answer a series of questions grounded in an image. |
| Approach: | They propose dual attention networks (DAN) for visual reference resolution in VisDial. |
| Outcome: | The proposed model outperforms the previous state-of-the-art model on VisDial datasets. |
Copied to clipboard
| Challenge: | Existing prompt tuning methods tend to learn spurious or entangled representations, leading to poor generalization to unseen concepts. |
| Approach: | They propose a prompt tuning technique that tunes the learnable prompt for pre-trained vision and language models. |
| Outcome: | The proposed method improves few-shot performance on vision and language tasks over existing prompt tuning methods. |
Copied to clipboard
| Challenge: | Existing visual question answering datasets assume only one ground truth answer for each question. |
| Approach: | They propose alternative answer sets (AAS) of ground-truth answers to address this limitation . they modify top VQA solvers to support multiple plausible answers for a question . |
| Outcome: | The proposed approach improves on the GQA dataset and shows that it is more efficient than previous approaches. |
Copied to clipboard
| Challenge: | a flurry of research has been conducted on the performance of state-of-the-art (SoTA) Vision Language Models (VLMs) on a variety of tasks. |
| Approach: | They propose a benchmarking tool to analyze performance of SoTA Vision Language Models (VLMs) on three tasks: Question Rephrasing, Image Restyling, and Context Reasoning. |
| Outcome: | The proposed model achieves absolute improvements of 5.7% and 12.5% on widely used VLMs such as BLIP-2 and LLaVa 1.5M in terms of consistency over their existing counterparts. |
Copied to clipboard
| Challenge: | Large multi-modal models (LMMs) are capable of visual question answering (VQA) with unprecedented accuracy. |
| Approach: | They propose a calibration score that can be used to quantify uncertainty in visual question answering models. |
| Outcome: | The proposed calibration score is better calibrated than in text-only models for in-context learning. |
Copied to clipboard
| Challenge: | Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans. |
| Approach: | They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
| Outcome: | The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
Copied to clipboard
| Challenge: | Existing studies focus on generating QADs from image and question, but a novel task is needed to generate meaningful questions, correct answers, and challenging distractors. |
| Approach: | They propose a task to generate QADs from images and encode images together . they use contrastive learning to ensure consistency of QAD generated and tested . |
| Outcome: | Empirical evaluations on the benchmark dataset validate the performance of the proposed task. |
Copied to clipboard
| Challenge: | Recent advances in Vision-Language Models and the scarcity of high-quality multi-modal alignment data have inspired numerous researches on synthetic VLM data generation. |
| Approach: | They propose a multi-modal data construction pipeline that organizes the final output into a Python code format. |
| Outcome: | The proposed pipeline improves visual question answering and visual grounding benchmarks across different VLMs. |
Copied to clipboard
| Challenge: | Recent work has shown that datasets contain incidental correlations created by idiosyncrasies in the data collection process. |
| Approach: | They propose a method that detects and ignores dataset-specific correlations by introducing a new method that makes them conditionally independent. |
| Outcome: | The proposed method detects and ignores these kinds of dataset-specific correlations, and does not require the bias to be known in advance. |
Copied to clipboard
| Challenge: | Existing approaches to question generation for interactive retrieval have constrained answer spaces, limiting the amount of information a model can gain in a single turn. |
| Approach: | They propose a method that incorporates presupposition handling into question selection and belief updates. |
| Outcome: | The proposed method increases accuracy over the past state-of-the-art by 14% while resulting in 48% more efficient games in human evaluations. |
Copied to clipboard
| Challenge: | Manga is a richly multimodal narrative form that blends images and text in complex ways. |
| Approach: | They propose two benchmarks for multimodal manga understanding: mangaOCR and mangaVQA . mangaVQ consists of 526 high-quality, manually constructed question-answer pairs . |
| Outcome: | The proposed model is finetuned from the open-source LMM Qwen2.5-VL . it compares with proprietary models such as GPT-4o and Gemini 2.5 to evaluate its performance . |
Copied to clipboard
| Challenge: | Existing instruction-tuned models struggle to adhere to a query with multiple intentions, which impairs their performance when the completion of several tasks is demanded by a single command. |
| Approach: | They develop an automatic process that turns existing data into diverse and complex task chains and a new benchmark to evaluate a model’s ability to follow all the instructions in a sequence. |
| Outcome: | The proposed model can follow instructions better and deliver higher results in coding, maths, and open-ended generation. |
Copied to clipboard
| Challenge: | Existing research on visual question answering is limited to information explicitly present in an image or a video. |
| Approach: | They propose a vision-language question answering task based on a CLEVR dataset . they modify existing methods and propose baseline solvers for this task . |
| Outcome: | The proposed model motivates the development of better vision-language models . it provides insights about the capability of diverse architectures to perform joint reasoning over image-text modality. |
Copied to clipboard
| Challenge: | Existing models for low-resource languages with catastrophic forgetting pose several challenges, including learning to model multi-lingual scenarios. |
| Approach: | They propose to employ a continual learning strategy using parts-of-speech code-switching and replay adapter strategies to mitigate catastrophic forgetting gap while training LLM from LLM. |
| Outcome: | The proposed architecture is able to train LLMs from LLM and mitigate catastrophic forgetting gap on vision language tasks. |
Copied to clipboard
| Challenge: | iParaphrasing extracts visually grounded paraphrases, which are different phrasal expressions describing the same visual concept in an image. |
| Approach: | They propose a task to extract visually grounded paraphrases from images . they propose to model the similarity between the extracted VGPs using existing methods . |
| Outcome: | The proposed task extracts visually grounded paraphrases from images . the proposed method has the potential to improve multimodal language and image tasks . |
Copied to clipboard
| Challenge: | Attention mechanism has been used in Vision-and-Language (VL) tasks to bridge the semantic gap between visual and textual clues. |
| Approach: | They conduct a comprehensive analysis on understanding the role of attention alignment by looking into attention score calculation methods and checking how it represents the visual region’s and textual token’s significance for the global assessment. |
| Outcome: | The attention score calculation methods represent visual region’s and textual token’s significance for the global assessment. |
Copied to clipboard
| Challenge: | Existing approaches typically decompose only language queries, treating images as monolithic inputs. |
| Approach: | They propose a framework that decomposes both images and questions into visual sub-domains with corresponding sub-questions. |
| Outcome: | REDI achieves absolute accuracy improvements of 8.9%, 8.2%, and 16.0% over existing models. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. |
| Approach: | They propose large vision-Language Models to augment LLMs with visual inputs. |
| Outcome: | The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat. |
Copied to clipboard
| Challenge: | Recent studies have employed machine translation systems for cross-lingual VQA tasks . however, translated texts contain unique characteristics distinct from human-written ones, referred to as translation artifacts. |
| Approach: | They propose a machine translation system that can train models in multiple languages . they propose augmentation strategies that reduce translation artifacts in translated texts . |
| Outcome: | The proposed approach reduces translation artifacts in models across languages and languages. |
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |
Copied to clipboard
| Challenge: | Empirical results show that paragraph captions help answer more visual questions . |
| Approach: | They propose a visual and textual question answering model which uses paragraph captions as input . they use cross-attention to extract related information, then consensus to fuse the inputs . |
| Outcome: | Empirical results show that paragraph captions help answer more visual questions . the proposed model significantly improves the baseline model . |
Copied to clipboard
| Challenge: | Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world. |
| Approach: | They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts. |
| Outcome: | The proposed model outperforms the text-only variants on a commonsense question answering task. |
Copied to clipboard
| Challenge: | Existing methods focus on single-hop, single-modality, or short texts, limiting real-world applications . despite advances in visual question answering, this multihop setting remains underexplored due to a lack of quality datasets. |
| Approach: | They propose a framework for creating a high-quality dataset for multimodal multihop question answering . they use a 5-stage pipeline to acquire relevant multimodal documents from Wikipedia . |
| Outcome: | The proposed framework outperforms existing methods on multimodal multihop question answering datasets. |
Copied to clipboard
| Challenge: | Vision-language models struggle with spatial reasoning, a skill that humans excel at. |
| Approach: | They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics. |
| Outcome: | The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks. |
Copied to clipboard
| Challenge: | Existing approaches to fine-tune visual-language understanding (VLU) require tasks-specific designs and sufficient training data. |
| Approach: | They propose a simple yet efficient paradigm for low-resource Visual Language Understanding (VLU) they reformulate a series of VLU tasks as an open-book affinity-matching problem. |
| Outcome: | The proposed framework outperforms baselines in low-resource settings. |
Copied to clipboard
| Challenge: | Recent Vision and Language models have shown impressive performance across benchmarks . however, frontier models lack cultural awareness and can affect global cultural diversity . |
| Approach: | They propose a visual question answering benchmark to probe the knowledge of culture-specific concepts and evaluate the capacity for cultural adaptation through contextual information. |
| Outcome: | The proposed model shows large performance disparities between culture-specific and common concepts in the parametric setting. |
Copied to clipboard
| Challenge: | Existing evaluation datasets for external knowledge-based VQA lack a capability to determine which passage is useful for answering queries. |
| Approach: | They propose a visual question answering benchmark for vision language models based on retrieval augmented generation (RAG) the proposed benchmark includes five input passages, a capability lacking in previous research. |
| Outcome: | The proposed benchmark includes five input passages and is validated using the state-of-the-art Llama3-based VLM, the Llava-Llamama-3 model. |
Copied to clipboard
| Challenge: | Multimodal research has picked up significantly in the space of question answering with the task being extended to visual question answering, charts question answering as well as multimodal input question answering. |
| Approach: | They propose a multimodal question-answering task that produces a unimodal textual output as the answer through human experiments. |
| Outcome: | The proposed framework outperforms existing frameworks on both automatic and human metrics. |
Copied to clipboard
| Challenge: | Previously, CLIP was only regarded as a powerful visual encoder. |
| Approach: | They propose a parameter-efficient fine-tuning strategy to boost CLIP's few-shot performance on a visual entailment task without introducing any additional pre-training procedure. |
| Outcome: | The proposed strategy achieves competitive zero/few-shot results on visual question answering and visual entailment tasks without introducing any additional pre-training procedure. |
Copied to clipboard
| Challenge: | Visual Programming is an alternative to end-to-end black-box visual reasoning models. |
| Approach: | They propose a visual programming strategy that leverages Large Language Models to generate the logic of a program in the form of its source code. |
| Outcome: | The proposed method improves ViperGPT on visual question answering and referring expression comprehension with an LLM. |
Copied to clipboard
| Challenge: | Existing pre-trained vision-language models suffer from inefficiency and linguistic signal overwhelmed by long visual sequences in cross-modal alignment. |
| Approach: | They propose a vision-language foundation model with cross-modal skip-connections that can be pre-trained end-to-end on large-scale image-text pairs with both discriminative and generative objectives. |
| Outcome: | The proposed model achieves state-of-the-art results on a wide range of vision-language downstream tasks, including image captioning, image-text retrieval, visual grounding and visual question answering. |
Copied to clipboard
| Challenge: | Existing datasets and models fail to consider critical aspects of medical diagnostics, authors argue . MMXU enables multi-image questions incorporating both current and historical patient data. |
| Approach: | They propose a dataset for MedVQA that focuses on identifying changes in specific regions between two patient visits. |
| Outcome: | The proposed dataset improves diagnostic accuracy by 20% by integrating historical data. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. |
| Approach: | They introduce DivScene, a large-scale dataset with 4,614 houses across 81 scene types and 5,707 kinds of target objects. |
| Outcome: | The proposed dataset provides a much greater diversity of target objects and scene types than existing datasets, enabling a comprehensive task evaluation. |
Copied to clipboard
| Challenge: | Multimodal semantic comprehension has attracted increasing research interest recently such as visual question answering and caption generation. |
| Approach: | They propose to use a large-scale multimodal instructional video dataset to support fine-grained comprehension research in specific domain. |
| Outcome: | The proposed dataset contains 2,800 videos from YouTube, spanning more than 420 hours in total. |
Copied to clipboard
| Challenge: | Existing multimodal tasks allow machines to understand images by describing or being asked in natural language. |
| Approach: | They propose a task that predicts the positions of images in a given document . they use a dataset of 66K multimodal documents with 320K images from Wikipedia . |
| Outcome: | The proposed task outperforms baselines while the performance is far from human. |
Copied to clipboard
| Challenge: | a systematicity gap exists between neural networks generalizing to new combinations of familiar concepts . conventionally trained neural networks struggle to generalize systematically . |
| Approach: | They propose to train a visual question answering model with CLEVR-HOPE as a diagnostic dataset to test this hypothesis. |
| Outcome: | The systematicity gap is reduced by increasing the diversity of training data, the authors show . the authors suggest that the more distinct attribute type combinations are seen during training, the more systematic the model will be. |
Copied to clipboard
| Challenge: | Currently, language-equipped vision systems such as VizWiz, TapTapSee, BeMyEyes, and CamFind are actively being deployed across a broad spectrum of users. |
| Approach: | They propose to identify collective outliers in active learning methods that are hard and often impossible for models to learn . they also propose to use visual inputs to identify these outlier examples as examples assigned low model confidence and prediction variability during training. |
| Outcome: | The proposed methods outperform random selection on visual question answering tasks. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks focus on single-chart tasks, neglecting multi-hop reasoning required to extract and integrate information from multiple charts. |
| Approach: | They propose a benchmark that evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
| Outcome: | The proposed benchmark evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
Copied to clipboard
| Challenge: | Existing studies find VH instances only in existing image datasets, which results in biased understanding of MLLMs’ performance under VH. |
| Approach: | They propose a tool called VHTest to generate a diverse set of VH instances from existing image datasets and a text-to-image generative model to generate VH images based on the text descriptions. |
| Outcome: | The proposed tool finds VH instances in existing image datasets and generates images based on the text descriptions. |
Copied to clipboard
| Challenge: | Visual Language Models have demonstrated remarkable capabilities across various tasks, including visual question answering and image captioning. |
| Approach: | They propose an end-to-end multimodal model that leverages speech instructions for reasoning-based visual question answering. |
| Outcome: | The proposed model can process and explain visual scenes from spoken input, moving beyond simple object recognition to reasoning-based interactions. |
Copied to clipboard
| Challenge: | Existing methods for reducing hallucinations incur a significant increase in latency. |
| Approach: | They propose a task-agnostic attention-guided head suppression strategy that can be seamlessly integrated during inference without incurring significant compute or latency overhead. |
| Outcome: | The proposed approach reduces hallucinations by 2.7x while maintaining F1 and improves throughput by 1.8% compared to existing methods. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a new challenge for AI. |
| Approach: | They propose a graph neural network architecture based on the recently proposed Graph Network (GN) . they generate visual features and encoded captions for an image to generate two GNs . |
| Outcome: | The proposed model rivals the state-of-the-art models on Visual7W, VQA-v2.0, and CLEVR datasets. |
Copied to clipboard
| Challenge: | Existing models for visual question answering are limited to the English language. |
| Approach: | They present a multimodal dataset for visual question answering tasks in the Hausa language. |
| Outcome: | The proposed dataset provides 12,044 gold standard English-Hausa parallel sentences that are semantically identical to the corresponding visual information. |
Copied to clipboard
| Challenge: | avrahami et al., 2022b,a): natural language instructions are often underspecified, requiring models to uncover their implicit meaning. |
| Approach: | They propose to use paired data to model the implicit meaning of instructions . they also propose to ground the model to localize where the edit has to be performed . |
| Outcome: | The proposed model performs better than state-of-the-art baselines on paired data, showing improvements in quality and faithfulness. |
Copied to clipboard
| Challenge: | RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior. |
| Approach: | They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests. |
| Outcome: | The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering. |
Copied to clipboard
| Challenge: | Existing approaches to learn and reason over language and vision data for downstream tasks such as visual question answering (VQA) and natural language for visual reasoning (NLVR) |
| Approach: | They propose a cross-modality relevance module that is used in an end-to-end framework to learn the relevance representation between components of various input modalities under supervision of a target task. |
| Outcome: | The proposed approach shows competitive performance on two different language and vision tasks using public benchmarks and improves the state-of-the-art published results. |
Copied to clipboard
| Challenge: | a traditional image captioning task uses generic reference captions to provide textual information about images. |
| Approach: | They propose a task that uses question-answer pairs to provide visual information instead of generic reference captions. |
| Outcome: | The proposed captioning with a purpose task can be tailored to meet user needs . question-answer pairs are used as a source of supervision for learning visual information needs a new task is proposed . |
Copied to clipboard
| Challenge: | Recent work has adapted vision-and-language models to generative tasks like image captioning. |
| Approach: | They propose an extension to LXMERT with training refinements to generate images from text. |
| Outcome: | The proposed model can generate images from pieces of text while still being comparable to existing models. |
Copied to clipboard
| Challenge: | Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content. |
| Approach: | They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions. |
| Outcome: | The proposed dataset contains 1000 images and 8,949 question-answer pairs. |
Copied to clipboard
| Challenge: | BiMediX2 is a bilingual (Arabic-English) large multimodal model that supports text-based and image-based medical interactions. |
| Approach: | They introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. |
| Outcome: | The model outperforms existing models by over 9% in English and more than 20% in Arabic evaluations. |
Copied to clipboard
| Challenge: | Large language and multimodal models have shown remarkable success on various benchmarks focused on specific skills such as general-purpose programming, math word problem-solving, and visual question answering. |
| Approach: | They propose a program synthesis benchmark based on real-world programming tasks . they propose 'fine-tuning pipeline' to boost performance of large language models . |
| Outcome: | The proposed model outperforms existing models on tasks that require a combination of skills on visual programming and programming. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks focus on static evaluation of large multimodal models . existing evaluation paradigms neglect a critical aspect of clinical practice: longitudinal analysis . |
| Approach: | They propose a temporal perception and reasoning benchmark to assess models' temporal grounding and consistency. |
| Outcome: | ELTLM features a hierarchical task taxonomy comprising Temporal Perception QA and Temporal Reasoning QA. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer. |
| Approach: | They propose a multimodal framework that leverages language guidance to answer questions more accurately. |
| Outcome: | The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models. |
Copied to clipboard
| Challenge: | Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites. |
| Approach: | They propose a large-scale document-based QA dataset that requires both visual and textual information to answer questions. |
| Outcome: | The proposed dataset incorporates multiple categories of questions and unanswerable questions from the document for realistic question-answering applications. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) gain expertise across diverse domains and modalities, a new study shows . scalable oversight becomes challenging when their capabilities surpass human evaluators. |
| Approach: | a new study extends the debate paradigm to a multimodal setting . it explores the potential for blind models to supervise and enhance the performance of sighted ones. |
| Outcome: | The proposed framework outperforms individual LLMs on multimodal tasks . it allows blind models to supervise and enhance the performance of sighted models . |
Copied to clipboard
| Challenge: | Existing methods to fine tune language agents with reasoning-action trajectories require high-quality model-generated samples, which are hard to obtain for challenging language agent tasks. |
| Approach: | They propose a method to employ reflection during inference without ground-truth feedback to improve agents more autonomously. |
| Outcome: | The proposed method improves self-training performance on open-source language agents by 7.6% and 14.1% respectively. |
Copied to clipboard
| Challenge: | Existing studies have focused mainly on visual–textual misalignment, leaving largely unexplored the MLLMs’ ability to preserve an original correct answer when confronted with misleading information. |
| Approach: | They propose a two-stage evaluation pipeline to quantify the response uncertainty phenomenon by eliciting each model’s original response on unperturbed inputs and injecting explicit (false-answer hints) and implicit (contextual contradictions) misleading instructions. |
| Outcome: | The proposed model overturns a correct answer in 65% of cases after receiving a single deceptive cue. |
Copied to clipboard
| Challenge: | Pre-trained vision and language models have demonstrated state-of-the-art capabilities over existing tasks involving images and texts. |
| Approach: | They analyze a visual question answering dataset tailored for info-seeking questions . they show that pre-trained visual and language models can use fine-grained knowledge . |
| Outcome: | The proposed dataset elicits models to use fine-grained knowledge learned during pre-training. |
Copied to clipboard
| Challenge: | Existing studies link hallucination to data or representation biases, but their causal origins remain unclear. |
| Approach: | They propose a causal framework to analyze and mitigate hallucination in vision-language models by using counterfactual analysis to estimate the Natural Direct Effect (NDE) of each modality and their interaction. |
| Outcome: | The proposed framework significantly reduces hallucination while preserving task performance while retaining reliability. |
Copied to clipboard
| Challenge: | Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time. |
| Approach: | They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization . |
| Outcome: | The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass . |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. |
| Approach: | They propose three confidence-based methods to enhance LVLMs' perception . they propose probabilistic and consistency-based signals are more reliable indicators . |
| Outcome: | Experiments on three LVLMs across three VQA datasets show that LVLs possess a reasonable perception level but there is room for improvement. |
Copied to clipboard
| Challenge: | Existing knowledge-aware text-based visual question answering methods are based on textual entities in images. |
| Approach: | They propose a visual text entity linking module that harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to perform visual text-entity linking. |
| Outcome: | The proposed approach surpasses the previous best approach by 23.3% on an absolute scale and establishes a new state of the art. |
Copied to clipboard
| Challenge: | Existing symbolic parsers lack flexibility to operate in complex, dynamic environments. |
| Approach: | They propose a framework that combines frame semantics with perceptual grounding to enable robots to interpret commands via multimodal logical forms. |
| Outcome: | The proposed framework produces over 11,000 image-command pairs and lowers the cost of manual parsers. |
Copied to clipboard
| Challenge: | Existing methods to train pretrained language models for zero-shot crossmodal tasks require crossmodal pretraining. |
| Approach: | They propose to inject visual concepts into the input text embedding space of a pretrained language model and build adaptation layers based on the intermediate representation of concepts. |
| Outcome: | The proposed model performs zero-shot crossmodal tasks without crossmodal pretraining . it is based on the injection of visual concepts as input tokens and augmentation in intermediate features . the proposed model achieves competitive or even better results in zero- shot and fine-tuning settings . |
Copied to clipboard
| Challenge: | Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions. |
| Approach: | They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding. |
| Outcome: | The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information. |
Copied to clipboard
| Challenge: | Recent training-free methods suggest that accuracy can be improved without fine-tuning. |
| Approach: | They propose a method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization. |
| Outcome: | The proposed method improves retrieval metrics for all contrastive models and datasets and does not require training on the reference database. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models perform well on visual question answering tasks, but it remains unclear whether their reasoning relies more on memorized world knowledge or on visual information present in the input image. |
| Approach: | They propose a dataset of visual-realistic counterfactuals that put world knowledge priors into conflict with visual input. |
| Outcome: | The proposed dataset puts world knowledge priors into conflict with visual input . it shows that model predictions shift toward visual evidence in mid-to-late layers . |
Copied to clipboard
| Challenge: | Vision-Language Models have shown impressive capabilities and notable failures in data visualization understanding tasks. |
| Approach: | They propose a benchmark to analyze how specific properties within a visualization type affect VLM performance. |
| Outcome: | The proposed benchmark examines how specific properties affect VLM performance . it shows that models exhibit steep drops on multi-hop reasoning and extraction errors increase with edge density . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on specific aspects of web tasks but lack comprehensive coverage. |
| Approach: | They propose a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. |
| Outcome: | The proposed model performs well on basic information extraction, but struggles with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. |
Copied to clipboard
| Challenge: | Existing approaches to learn multi-modal tasks are based on chain-of-thought . however, human thought processes are non-linear and employ dynamic adjustment and updating mechanisms. |
| Approach: | They propose a chain-of-thought technique that adjusts the length of the chain to improve the performance of generated prompts. |
| Outcome: | The proposed model improves multi-modal representation learning in visual, visual, and audio-visual tasks and also has good domain generalization performance due to better reasoning. |
Copied to clipboard
| Challenge: | Large vision-language models have shown impressive ability in various language tasks, especially with their emergent in-context learning capability. |
| Approach: | They propose a causal reasoning benchmark for multi-modal in-context learning from large vision-language models that incorporates visual inputs. |
| Outcome: | The proposed model outperforms existing models on three visual causal reasoning tasks and demonstrates their strengths and weaknesses. |
Copied to clipboard
| Challenge: | Existing datasets address understanding and generation in isolation, limiting the performance of unified vision large language models. |
| Approach: | They propose a dataset that facilitates mutual enhancement between multimodal understanding and generation. |
| Outcome: | The proposed framework integrates diverse visual and textual inputs and outputs, enabling comprehensive cross-modal reasoning and precise text-to-image alignment. |
Copied to clipboard
| Challenge: | Visual text compression is emerging paradigm for rendering text as images for processing by vision-language models. |
| Approach: | They propose a benchmark to assess VLM robustness under dense visual inputs. |
| Outcome: | Evaluating 13 general-purpose VLMs and 3 OCR-specialized models reveals performance drops sharply under increased density or reduced resolution; cross-task transfer between OCR, NIAH, and VQA is limited; and VQ is comparatively robust because low-level details are lost before high-level semantics. |
Copied to clipboard
| Challenge: | Existing mitigation approaches reduce hallucinated object mentions at the cost of degraded generation quality or require expensive retraining and task-specific supervision. |
| Approach: | They propose a lightweight framework for low-hallucination vision–language generation . it uses evidence-bounded minimal editing to revise or suppress unsupported referenced entities . |
| Outcome: | The proposed framework reduces hallucinations while maintaining or improving quality metrics. |